Health Data
"Health data" covers several genuinely different things. Confusing them is the most common cause of dashboards that answer no one's question.
Kinds of health data
Individual-level clinical data
Encounters, diagnoses, observations, medications, procedures. Generated during care, structured around the patient. Sensitive by default.
Aggregate routine data
Counts by facility and period — the classic HMIS model. Cheap to transmit, adequate for coverage monitoring, unable to answer questions about individuals or continuity of care. See DHIS2.
Registry data
Master lists: clients, facilities, health workers, products. Boring and foundational — most exchange failures are registry failures. See OpenHIE.
Survey and census data
Population-representative, periodic, expensive. The denominator that routine data usually lacks.
Surveillance data
Timely, targeted, often incomplete by design; optimised for detecting change rather than measuring level.
Logistics and financial data
Stock, consumption, claims, expenditure. Frequently the most reliable data in a health system, because money is reconciled.
Patient-generated data
From apps, wearables and self-report. High volume, variable quality, different consent basis.
Data quality dimensions
| Dimension | Question |
|---|---|
| Completeness | Did every reporting unit report? |
| Timeliness | Did it arrive when it was needed? |
| Accuracy | Does it match the source record? |
| Consistency | Do related values agree? |
| Validity | Is it within plausible ranges and code sets? |
| Uniqueness | Is the same person or event counted once? |
Quality is produced at the point of collection. Downstream cleaning can detect problems; it cannot create information that was never recorded. The strongest lever is making the data useful to the person entering it.
Denominators
Routine data supplies numerators. Coverage indicators need denominators — population estimates, target populations, catchment areas — and these are usually the weakest part of any coverage figure. Report the denominator source alongside the indicator, or the number is uninterpretable.
De-identification
Removing names is not de-identification. Re-identification risk comes from combinations: date of birth, sex, facility and visit date can identify a person in a small population.
Techniques: removing direct identifiers, date shifting, generalising geography and age, aggregating small cells, and formal approaches such as k-anonymity or differential privacy for release.
Small numbers are the recurring hazard in health reporting — a cell of one in a district table can be a disclosure. Suppression rules should be set before publication, not after a complaint.
Related
- Open data — publishing health data responsibly
- Data governance — who decides
- ISO 27799 — protecting personal health information
- GDPR — legal obligations